Genome Research
● Cold Spring Harbor Laboratory
Preprints posted in the last 90 days, ranked by how well they match Genome Research's content profile, based on 468 papers previously published here. The average preprint has a 0.27% match score for this journal, so anything above that is already an above-average fit.
Shin, D.-H.; Jeon, J.; Joe, S.; Jeon, Y.; Yang, J. O.; Bhak, J.; Baek, S. A.; Byun, G.; Shin, E.-S.; Kwon, Y.; Choi, H.-J.; Kim, J.-H.; Haam, K.; Yoo, J.; Song, K. J.; Mok, J.; Jeon, S.; Jeong, H.; Bhak, J.
Show abstract
Here, we present the first graph-based Korean Pangenome Reference (K-PanRef), constructed from 14 healthy Korean individuals. K-PanRef comprises 13 high-quality diploid Korean genome assemblies (mean QV ~62.0) and KOREF1-G-TTAGGA, the first complete Korean reference genome. Integration of these assemblies generated a ~3.2-Gb pangenome graph containing ~39.3 million nodes and ~53.8 million edges, with the accumulation of common sequences (frequency [≥]10%) reaching a plateau. Additionally, K-PanRef contains ~4.3 million Korean-specific small variants and ~76.0 thousand Korean-specific SVs absent from the Chinese and human pangenome references, improving the representation of Korean genetic diversity relative to these references. To evaluate its utility for short-read-based SV analysis, we genotyped 75 whole-genome sequencing (WGS) samples, including 15 patients with early-onset myocardial infarction (MI). Although constructed entirely from healthy genomes, K-PanRef supported the identification of putative disease-relevant SVs in this exploratory application. K-PanRef-based genotyping identified ~95.6 thousand small variants and 820 SVs observed only in the early-onset MI samples. Among the early-onset MI-group SVs, 491 were absent from public databases, suggesting that they may represent previously unrecognized candidate variants related to early-onset MI. Of these, 164 SVs overlapped 134 genes, of which 89 had reported associations with 42 cardiovascular diseases or traits, including eight genes previously linked to MI. Together, these results establish K-PanRef as a valuable resource for representing Korean genetic diversity and enabling more comprehensive discovery of population-specific and novel putative disease-relevant variants from short-read sequencing data.
Samano, A.; Chakraborty, M.
Show abstract
Centromeres ensure faithful chromosome segregation despite being embedded within rapidly evolving repetitive DNA, a contradiction known as the centromere paradox. While centromere identity is defined by the histone variant CENP-A, how conserved function is maintained amid rapid DNA turnover remains unclear. Here, we generate highly contiguous genome assemblies from single Drosophila melanogaster individuals that, for the first time, resolve a chromosome through its centromere, linking the chromosome 3 arms within a continuous sequence. Comparative assemblies from wild-derived strains reveal extensive structural variation in pericentromeric satellites, including large-scale expansions, contractions, and sequence divergence. Despite this variation, the CENP-A-associated centromeric core exhibits conserved organization across strains. Integration of Hi-C interaction maps with sequence analyses shows that flanking dodeca satellite arrays form a spatially interacting domain that bridges both sides of the centromere, whereas adjacent Prodsat arrays are more variable and show weaker interactions. These results support a model in which rapidly evolving centromeric DNA is constrained by conserved higher-order architecture, providing a framework for reconciling the rapid evolution of centromere sequence with its conserved function.
Martini, J.; Williams, R. A.; Smith, R. G.; Zhou, Y.; Liu, Y.
Show abstract
The coordinated activities of histone modifications and chromatin-associated proteins establish chromatin states that regulate genome function and cellular identity. However, the organizational principles that distinguish chromatin states across cell types remain incompletely understood. Here, we integrated genome-wide profiles of CTCF, H3K27ac, H3K9ac, H3K27me3, and H3K9me3 at 1-kb resolution to generate a unified representation of chromatin organization in human cells. Unsupervised embedding resolved five principal chromatin states corresponding to constitutive heterochromatin, CTCF-associated architectural chromatin, transcriptionally active chromatin, mixed repressive chromatin, and Polycomb-associated chromatin. Comparative analyses of HCT116 and K562 cells revealed that cell-type-specific epigenomic differences arise predominantly through remodeling of Polycomb-associated chromatin, whereas the remaining chromatin states exhibit broadly similar epigenomic signatures and comparatively limited remodeling. Consistent with this observation, principal component analysis identified H3K27me3 as the primary contributor to genome-wide epigenomic divergence, whereas CTCF represented a secondary contributor. Integration with Hi-C data further demonstrated that CTCF-associated chromatin is strongly enriched at chromatin loop anchors and other local architectural features, linking chromatin-state organization to three-dimensional genome architecture. Together, our findings identify distinct chromatin-state classes that organize the human epigenome and reveal Polycomb-associated chromatin as a key determinant of cell-type-specific epigenomic differences.
Delgado, A. A.; Samano, A.; Chakraborty, M.
Show abstract
The maintenance of functional repeat arrays on nonrecombining sex chromosomes presents an evolutionary paradox: tandem repeats are intrinsically unstable yet must preserve sequence identity and copy number to remain functional. The Y-linked Suppressor of Stellate (Su(Ste)) locus in Drosophila melanogaster is a large tandem array that produces piRNAs to silence the X-linked meiotic driver Stellate, but how such arrays are maintained remains unclear. Here, we reconstruct and compare repeat-resolved assemblies of the Su(Ste)/PCKR tandem array across three strains and show that the array is partitioned into discrete domains of elevated sequence identity. These domains exhibit an alternating pattern of similarity, in which nonadjacent regions are more similar to each other than to neighboring regions, and this organization is conserved across strains. Copy-number variation occurs primarily within specific domains, while overall array architecture remains stable. These results indicate that concerted evolution in the Su(Ste) array operates within structurally defined domains rather than uniformly across the array. The association of domain boundaries with inverted repeat elements suggests that higher-order structure constrains gene conversion, shaping both sequence homogenization and copy-number dynamics. In contrast, Y-linked rDNA arrays show uniform sequence similarity across long genomic distances, indicating a distinct mode of homogenization. Together, our findings demonstrate that gene conversion on nonrecombining chromosomes is structured by higher-order array architecture, providing a general framework for the maintenance of functional repeat arrays.
Mohanty, S. K.; Marin, M. G.; Smeds, L.; Chiaromonte, F.; Huber, C. D.; Makova, K. D.; Human Pangenome Reference Consortium,
Show abstract
G-quadruplexes (G4s), non-canonical DNA structures whose sequence motifs occupy approximately 1% of the human genome, are important for myriad cellular functions, including regulating transcription and replication. Yet they also contribute to genomic instability by increasing mutations and structural variation. Despite their significance, G4 motifs have not been studied in detail across multiple human genomes. Here, we conducted a comprehensive analysis of presence/absence and sequence variation, measured selection strength, and evaluated gene expression regulation potential for predicted G4s (pG4s) across population groups in the second release of the Human Pangenome Reference Consortium dataset, comprising high-quality, near-telomere-to-telomere diploid genomes from 231 individuals worldwide, along with three reference assemblies. Across the human pangenome, we identified over 353 million pG4s, including 1.15 million pG4s absent from reference assemblies but shared across other haplotypes. Our analysis revealed that pG4 sharing patterns recapitulate human population structure: African individuals displayed lower levels of pG4 sharing than non-Africans, whereas East Asian individuals exhibited higher levels of sharing. By analyzing the site frequency spectrum across various genomic annotations, we computed and compared selection coefficients (Sd) at pG4 vs. non-pG4 sites. As expected, the strongest purifying selection (Sd [≥] 10) was detected at protein-coding exons, where pG4 sites had similar or lower selection coefficients compared with those for pG4 sites. Strikingly, this pattern reversed at regulatory regions: although purifying selection was weaker overall at promoters, introns, enhancers, and replication origins (1 [≤] Sd < 10), pG4 sites at these regions experienced stronger selection than non-pG4 sites--suggesting that pG4s play functional roles outside coding sequences. Additionally, by integrating pG4 data with long-read transcriptome data profiles from this large cohort, we found that pG4s located at promoters and at (or near) exon-intron junctions may influence variation in gene expression levels and transcript isoforms, respectively, across the human pangenome individuals. Leveraging extensive population-scale data, our research illuminates the fundamental importance and functional relevance of G4s across human genomes.
Kim, H. J.; Kim, S. M.; Yoon, K.; Kim, B.-J.; Jin, H. J.; Kim, Y. J.
Show abstract
While long-read sequencing technologies (e.g., PacBio Revio, ONT) have revolutionized high-quality genome assembly for the human pangenome, mitochondrial genome (mtDNA) analysis still largely relies on short-read and Sanger sequencing. However, short-read sequencing often lacks the resolution required to resolve complex variations due to the unique features of mtDNA, such as high mutation rates and repetitive homopolymeric regions, which frequently lead to alignment artifacts and mapping ambiguities. To address this, we evaluated whether applying long-read sequencing to mtDNA improves analytical quality in empirical data. Through comprehensive bioinformatics analyses, we compared the performance of long-read sequencing against short-read sequencing and microarrays. Our results revealed that long-read sequencing detected the highest number of variants (n = 533), significantly outperforming both short-read sequencing (n = 525) and microarrays (n = 49). Notably, both sequencing methods provided significantly higher resolution in haplogroup assignment compared to microarrays in terms of phylogenetic depth (p < 0.05). Long-read sequencing demonstrated superior detection power, particularly for InDels. We identified two novel non-synonymous variants, including a unique InDel detected exclusively by long-read sequencing. Protein modeling and stability analysis validated that this InDel causes structural instability (RMSD > 2.0 [A],{Delta}{Delta} G = -45.21 kcal/mol). Furthermore, we confirmed that this novel InDel is shared among haplogroup A samples in both the 1000 Genomes Project ONT dataset and the Korean population, highlighting the practical implications of long-read sequencing for molecular biology and population genetics.
Likhite, M.; Andrews, G. R.; Gao, M.; Ramalingam, V.; Hecht, V.; Kundaje, A.; Moore, J. E.
Show abstract
Gene regulation depends on coordinated interactions between promoters and distal cis-regulatory elements, yet understanding how these regulatory elements communicate remains a fundamental challenge in mammalian genomics. Chromatin interaction assays provide one approach for identifying potential regulatory relationships, but interpreting the biological significance of individual interactions remains difficult; chromatin interactions comprise multiple biologically distinct classes that are only partially captured by any single assay. Here, we integrate Hi-C, RNAPII ChIA-PET, and CTCF ChIA-PET with the ENCODE Registry of candidate cis-regulatory elements (cCREs) and complementary functional genomic datasets to develop an integrative framework for classifying and interpreting promoter-centric chromatin interactions. Using this framework, we identify a distinct class of candidate architectural promoter-enhancer interactions that are characterized by increased recurrence across cellular contexts, broader promoter connectivity, and reduced dependence on linear genomic proximity. We further show that many regulatory elements anchoring these interactions transition between enhancer and CTCF-only states while maintaining stable chromatin interactions. These dual-state regulatory elements also acquire context-specific transcription factor inputs within evolutionarily conserved architectural scaffolds, suggesting that stable chromatin architecture can be repeatedly repurposed for new regulatory functions. Genes connected to these dual-state regulatory elements are enriched for developmental and signaling pathways and exhibit increased expression specificity across cell types, consistent with specialized roles in context-dependent gene regulation. Together, our findings provide a biologically informed framework for classifying and interpreting chromatin interactions and support a model in which conserved chromatin architecture provides a stable foundation upon which new regulatory programs evolve.
Yuen, Z. W. S.; Leeder, N.; Udumanne, T.; Garvie, A.; Wong, L.; Weiss, E.; van Loon, L.; Ganley, A.; Hannan, R.; Eyras, E.; Hein, N.
Show abstract
Ribosomal RNA (rRNA) provides the structural and catalytic core of ribosomes and is encoded by ribosomal RNA genes (rDNA) arranged in tandem repeat arrays. rDNA copy number (CN) is highly dynamic, representing a clinically relevant form of structural variation, but its accurate quantification has been challenging due to its highly repetitive and GC-rich nature. Here, we present RICO (Ribosomal DNA Integrated Copy Number and Methylation Analysis), a novel computational pipeline for integrated estimation of rDNA CN and methylation using nanopore long-read sequencing. RICO leverages long sequencing reads that span entire rDNA repeats, mapped to an rDNA-augmented reference genome, and normalizes coverage using an array of single-copy genes. We show that RICO provides accurate rDNA CN estimates in simulated datasets and reproducible measurements across human samples, with strong agreement to short-read sequencing and PCR-based methods. As biological validation, RICO detects a ~40% reduction in rDNA CN in Atrx-knockout mouse cells, consistent with established effects of ATRX loss on rDNA CN, and captures detected increased total and active rDNA CN in malignant cells from a MYC-driven B-cell lymphoma mouse model, in line with prior psoralen-based chromatin studies. Applying RICO to independent human cohorts, we uncover that individuals with higher total rDNA CN consistently exhibited higher fractions of high-methylated rDNA copies, suggesting a dosage compensation mechanism that potentially maintains a similar number of active rDNA copies across individuals. Together, RICO enables integrated analysis of rDNA CN and methylation state, providing a scalable framework for investigating rDNA regulation across population and disease studies.
Gunasekera, S.; Carlson, M.; Ray, M.; Larschan, E.
Show abstract
Nuclear bodies are nucleoprotein complexes with established functions that target chromatin at specific locations and regulate specific RNA processing functions, thereby influencing gene expression. However, the mechanisms that define how nuclear bodies are targeted to specific locations within the genome where they function remain poorly understood. One significant challenge is capturing and understanding the multiple cell-specific interactions occurring in these complexes, arising from RNA components interacting with each other and with DNA and nucleic acid-binding proteins within the context of the nucleuss three-dimensional organization. Mapping these interactions is critical for elucidating mechanisms such as RNA splicing, a key driver of cell-specific transcript diversity. Here, we use RNA-DNA Split Pool Recognition of Interactions by Tag Extension (RD-SPRITE) to characterize, for the first time, sex-specific RNA-RNA and RNA-DNA interactions in Drosophila S2 (male) and Kc (female) cells. We determined the sex-specific RNA-RNA interaction map within the nucleus and, using RNA-DNA interaction data, pinpointed the target loci of various RNA molecules, including small nuclear RNAs (snRNAs), which are core components of the spliceosome-a ribonucleoprotein complex involved in RNA splicing. Based on RNA-RNA interaction data, we also identified novel long non-coding RNAs that may regulate splicing. Furthermore, we investigated the role of transcription factor (TF) CLAMP in sex-specific targeting of the spliceosome. We generated RD-SPRITE datasets in the presence and absence of CLAMP, a key TF involved in dosage compensation, sex-specific RNA splicing, and chromatin organization. We determined that CLAMP regulates global changes in spliceosomal interactions with chromatin, inhibits aberrant snRNA interactions, and regulates sex-specific interactions of RNAs involved in splicing function. Additionally, our dataset provides a valuable resource for investigating additional processes, such as miRNA-mediated silencing, nucleolar functions of snoRNAs, and Cajal body functions of scaRNAs, among others. To facilitate broad community use, we have developed a computational platform, "FlySprite," that enables Drosophila researchers to explore sex-specific RNA-RNA interactions, as well as DNA targets of RNA clusters, through a user-friendly interface.
Carvalho, A. B.; Kim, B. Y.; Uno, F.
Show abstract
We recently showed that Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) have very strong sequencing bias against simple satellites, probably caused by single-stranded DNA folding into non-canonical (non-B) structures during sequencing. Here we extend these observations by computational and experimental approaches in the Drosophila and human genomes. We found that (i) only a small subset of simple satellites cause sequencing bias; many satellites (e.g., (ACTGGG)n) are benign and easily sequenced. (ii) The biases most likely are caused by two distinct non-B DNA structures: triplex DNA formed by some, but not all, AG-rich satellites (only those predicted to form strong mirror repeats), and hairpins formed by some, but not all, AT-rich satellites (only those predicted to form fairly strong inverted repeats). (iii) The correlation between the predicted stability of these non-B structures, and the strength of sequencing bias indicates that non-B DNA is indeed the culprit. (iv) The likely source of these non-B structures is single-stranded DNA formed during ONT and PacBio sequencing, and hence its removal might solve the bias. We tested this by adding single-strand binding protein to ONT sequencing, and found that it irreversibly kills the flow cells. (v) A recent sequencing effort in Drosophila melanogaster using very high depth Ultra-Long ONT sequencing (967x) still failed to assemble many genes located near satellite blocks. Brute-force will not solve the problem; instead, further investment is needed by the sequencing companies to achieve truly unbiased sequencing.
Mercuri, R. L. V.; Mombach, D. M.; dos Santos, F. R. C.; Perez-Schindler, J.; Huang, Y.; Spealman, P.; Pintacuda, G.; Al'Khafaji, A.; Donnard, E. R.; Claussnitzer, M.; Galante, P. A. F.
Show abstract
Transposable elements (TEs) not only account for half of the human genome sequence but also generate transcripts that contribute to transcriptomic diversity. Yet, their repetitive nature has hindered accurate quantification of the full TE-derived transcriptome, a challenge that long-read sequencing can overcome. Here, we combined multiplexed arrays isoform sequencing (MAS-ISO-seq) with a dedicated computational framework (TEscape) to perform an in-depth annotation of the human TE transcriptome. To capture the breadth of human transcriptome diversity, we profiled six representative cell types spanning three distinct biological contexts, including metabolism with, primary patient-derived adipogenic cells at two differentiation stages, and iPSC derived hepatic progenitor cells; the nervous system with iPSC-derived neurons, neural progenitor cells (NPCs), and pluripotency using induced pluripotent stem cells (iPSCs). Together, these datasets yielded over 235 million full-length long reads. First, to assess data coverage and transcriptome depth, we quantified protein-coding gene expression, detecting 14,312 genes (73.6% of all annotated protein-coding genes), which is a level consistent with deep and comprehensive transcriptome representation. Second, focusing on TE-derived transcripts, we identified >83,000 previously unannotated isoforms, the vast majority (84%) originating from a complex combination of multi-TEs. We also identified solo TEs, which are predominantly from LINE1 (14%). We confirmed that TE-transcripts are able to be exemplified by signatures detected in Liver Hepatocellular Carcinoma (LICH). Together, MAS-ISO-seq and TEscape establish the first long-read-based, high-resolution atlas of transcribed human TEs, providing a foundational resource for integrative transcriptome analyses and for investigating TE expression and regulation in health and disease. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=112 SRC="FIGDIR/small/737305v1_ufig1.gif" ALT="Figure 1"> View larger version (35K): org.highwire.dtl.DTLVardef@18158aeorg.highwire.dtl.DTLVardef@e51fdforg.highwire.dtl.DTLVardef@8f9504org.highwire.dtl.DTLVardef@804113_HPS_FORMAT_FIGEXP M_FIG Graphical Abstract C_FIG
Bao, W.; Qin, F.; Xiao, F.
Show abstract
Chromosomal copy number alterations (CNAs) are key drivers of tumor evolution, disease progression and therapeutic resistance, and the identification of them is an important step to delineate tumor clonal structure. However, accurately resolving CNA landscapes from single-cell data remains challenging. Most existing tools analyze one omics layer at a time and are susceptible to assay-specific noises, limiting their ability to recover shared or modality-specific CNAs. Recent single-cell multi-omics techniques enable joint sequencing of multiple molecular layers in the same cells, yet in silico methods that fully exploit such complementary multi-modal data for CNA analysis are still missing. Here we present a single-cell multi-omics integration framework, MintCNA, a unified framework for CNA detection from paired multi-omics data. MintCNA integrates traditional statistical modeling with embedded deep learning structure to enhance CNA profiling across multi-omics. We use an attention-guided convolutional autoencoder for data denoising and perform multivariate change-point detection utilizing a sliding-window screening and ranking procedure. Missingness-adjusted CUSUM statistics are constructed which jointly aggregate omics features by a data-adaptive projection to detect genome-wide chromosomal breakpoints. Across various simulations and applications to a colorectal cancer multi-omics dataset, MintCNA consistently outperforms existing single-omics CNA callers in detection accuracy. MintCNA provides a single-cell CNA tool that integrates paired scDNA-seq and scRNA-seq, supporting the study of intra-tumor heterogeneity and tumor evolution.
Mei, C.; Ness, J.; Nakai, K.; Wunderlich, Z.
Show abstract
Developmental processes depend on carefully coordinated gene expression. Expression is modulated by the binding of transcription factors (TFs) to cis-regulatory elements (CREs), like enhancers and promoters. Many computational and experimental approaches have been developed to find CREs, particularly enhancers, in the genome, each with strengths and caveats. Given the increasing availability of ATAC-seq data and methods to find TF binding therein, we hypothesized that we could use TF footprinting tools to find clusters of TF binding events within accessible chromatin that may act as CREs. Using Drosophila anterior-posterior patterning network as a test bed, we used a digital genomic footprinting tool (DGT), TOBIAS, on previously published early embryo ATAC-seq data to characterize the TF footprint landscape of 16 TFs essential for embryonic patterning. Even in this system, with its extensive enhancer annotation, most footprinted TF binding sites lie outside of known enhancers, with intergenic and intronic regions hosting the highest TF footprint count, albeit at low density. To find potential novel enhancers, we identified high-density TF footprint clusters that are highly conserved and overlap with active enhancer histone mark signals. Five high confidence candidates were selected for reporter assay validation and all five were found to drive spatially patterned expression in the embryo. This study shows that even in a highly characterized system, the analysis of footprinted TF binding sites in ATAC-seq data can uncover new regulatory regions and suggests this approach may be helpful in using existing ATAC-seq data to find novel CREs. ARTICLE SUMMARYGiven the increasing availability of ATAC-seq datasets, workflows to exploit the data to uncover new cis-regulatory elements (CREs), including enhancers, are valuable. Using early anterior-posterior patterning in the Drosophila embryo as a test case, we find that previously published transcription factor footprinting tools and ATAC-seq data can be analyzed to yield new candidate CREs. Experimental validation confirms the activity of selected candidate CREs, suggesting that existing data can be analyzed to find novel regulatory elements.
Czapiewski, R.; Chiang, M.; Ding, J.; Naughton, C.; Grimes, G. R.; Marenduzzo, D.; Gilbert, N.
Show abstract
Cell-to-cell transcriptional heterogeneity, or noise, is an intrinsic property of the transcriptome with implications for development, disease progression, and aging. Bulk RNA-seq masks this variability by averaging gene expression across cells, whereas single-cell RNA sequencing (scRNA-seq) resolves it. Nevertheless, separating biological noise from technical variance remains challenging, particularly across platforms with different chemistries. We benchmarked two widely adopted technologies, Evercode WT (SPLiT-seq, Parse Biosciences) and Chromium (10x Genomics), on human lymphoblastoid nuclei. Evercode WT achieved targeted sequencing depth and nuclei number far more reliably, and its random-hexamer priming yielded more intronic reads and non-coding RNA genes; Chromium recovered more cells and detected polyadenylated transcripts and cell-line markers more sensitively. Despite these opposing biases, the platforms showed comparable gene detection and strongly correlated expression profiles. Using datasets from both platforms, we defined a noise metric detrended from mean expression and showed that per-gene estimates were reproducible across chemistries. Noise was lower in G2M than in G1 and was most strongly associated with gene length rather than exonic length. Expression of genes with CpG-island promoters was less variable than that of those without. This study establishes a platform-independent basis for quantifying transcriptional noise and a framework for selecting an appropriate scRNA-seq platform.
Chapis, M.; Manousi, D.; Diblasi, C.; Brekke, C.; Kwak, J.; Ponce De Leon, A. V.; Arnyasi, M.; Fenstad, R.; Boison, S.; Saitou, M.
Show abstract
Structural variants (SVs) can affect gene regulation, but they are difficult to include in expression genetic studies when large RNA-seq cohorts lack whole-genome sequencing. This is common in non-human and non-model systems, where whole-genome sequencing at population scale remains costly. As a result, expression quantitative trait locus (eQTL) studies often rely on single nucleotide polymorphism (SNP) markers. These analyses can identify expression-associated regions, but often provide limited biological interpretation of the underlying regulatory mechanisms. Here, we used Atlantic salmon as a study system to test whether graph-genotyped SVs can be imputed into a SNP-array-genotyped RNA-seq cohort and used to interpret regulatory haplotypes. SVs were discovered from two long-read-sequenced individuals, supplemented with short-read SV and SNP calls from a 112-individual whole-genome-sequenced reference panel, graph-genotyped, jointly phased with SNPs, and imputed into 906 offspring with gill RNA-seq and SNP-array genotypes. After size filtering, the imputed SV catalogue contained 100,269 variants and showed nonuniform genomic distributions associated with sex-specific recombination landscapes. Association testing identified 51 SV-eQTL candidates, including 35 cis and 16 trans associations. These candidates were enriched for short-read-derived variants, indicating that short-read supplementation can recover regulatory variants missed by small-scale long-read discovery. SV-eQTL candidates were more strongly tagged by nearby SNPs than non-associated variants generally, but individual SNP lead markers often failed to capture the same eQTL signals in conditional regression. Retained candidates after the conditional analysis included target-gene-overlapping deletions, nearby local variants without target-gene overlap, trans associations, and short insertions with opposite effects on gene expression. These results show that imputed graph-genotyped SVs can add biological interpretation to possible regulatory haplotypes.
YUAN, J.; XUE, Z.; TANG, H.; LIU, Y.; WANG, J.
Show abstract
Short tandem repeats (STRs) are a rich and highly polymorphic source of human genetic variation, but representing and genotyping them in pangenome graphs remains challenging. Explicitly encoding each STR allele as a separate graph path results in increasingly complex local structures as cohort diversity increases, leading to larger index sizes and requiring significant resources for graph reconstruction when new alleles are introduced. Here, we propose STR-PG, a topologically decoupled genome-wide framework that separates stable locus representation from scalable STR allele content. STR-PG uses topologically fixed pointer nodes to represent each target locus, while allele sequences, repeat counts, motif annotations, and population frequency metadata are stored in an external registry. Short reads are mapped to STR loci via syncmer-based flanking anchors, and genotyping is performed within a locus-specific candidate space using allele-level alignment likelihood and Bayesian inference. Newly supported alleles can be integrated through registry-level updates without the need to rebuild the graph structure. Evaluations using simulated whole-genome sequencing data, 1000 Genomes Project (1kGP) samples, and r real whole-exome sequencing data from matched whole-blood-cell controls demonstrate that STR-PG maintains accurate genotyping results across various STR classes, reproduces expected population structures, and substantially reduces the computational cost of integrating additional alleles. STR-PG provides a compact and scalable framework for population-scale STR analysis using short-read sequencing.
Rossini, O.; Cleynen, A.; Shirokikh, N. E.
Show abstract
Untranslated regions (UTRs) flanking the coding sequence govern mRNA translation, localisation, stability, and decay, making accurate UTR boundaries essential for quantitative RNA sequencing and the study of post-transcriptional control in Saccharomyces cerevisiae and beyond. Reference transcriptomes built from short-read sequencing have been invaluable to the yeast community, yet in a genome as gene-dense as that of S. cerevisiae, short reads frequently cannot be assigned to a single transcript of origin, leaving roughly one quarter of transcripts without a confidently defined UTR. Here we use Oxford Nanopore Direct RNA Sequencing (DRS), in which each full-length polyadenylated molecule is read end to end, to resolve this ambiguity and deliver two complementary resources. First, an updated, ready-to-use S. cerevisiae S288C reference: change-point segmentation of per-gene DRS coverage defined boundaries for 5,416 of the 6,695 annotated genes, and a merge-max rule retaining the longer UTR from each source ensures no gene loses existing annotation. The result adds previously absent UTRs to 927 (5') and 896 (3') genes and extends 29.4% of 5' and 26.1% of 3' boundaries among comparable genes. Second, the complete, documented pipeline so that any laboratory can rebuild or update a transcriptome from its own DRS data. Validation on two independent datasets shows improved mapping rates, reduced soft-clipping, and metagene profiles consistent with genuine transcript signal.
Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.
Show abstract
Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.
McGowan, J.; Lipscombe, J.; Kilias, E. S.; Barker, T.; Catchpole, L.; Durrant, A.; Irish, N.; McTaggart, S.; Warring, S. D.; Gharbi, K.; Richards, T. A.; Hall, N.; Swarbreck, D.
Show abstract
Multiple displacement amplification (MDA) enables whole-genome amplification from single cells, but introduces chimeric artifacts that severely compromise downstream analyses, particularly with long-read sequencing. Here, we systematically evaluate long-read PacBio HiFi sequencing of MDA amplified DNA from single cells using the model green alga Chlamydomonas reinhardtii. We show that MDA-derived libraries exhibit highly uneven coverage and extreme chimera rates impacting up to 70% of reads, leading to thousands of artefactual structural variants and misassemblies when assembled using algorithms designed for bulk sequencing. To overcome these challenges, we developed lrSAGA (long-read Single Amplified Genome Assembly), a novel tool to assemble long-read MDA sequencing datasets. Assemblies generated using lrSAGA are more complete, more contiguous, and have 75-95% fewer misassemblies compared to conventional assembly algorithms. Although overall contiguity is limited by MDA coverage dropouts, we demonstrate that up to 68% of the C. reinhardtii genome can be accurately assembled from just a single haploid cell. We further validated lrSAGA using published Oxford Nanopore and PacBio HiFi data from single or half Caenorhabditis elegans worms, generating accurate and highly complete assemblies. Applying our approach to single protist cells isolated from environmental water samples, we performed PacBio HiFi single-cell genome sequencing of four uncultivated microbial eukaryotes: an amoeboflagellate from the Naegleria genus, a flagellate from the Bodo genus, and two deep-branching flagellates from the enigmatic CRuMs supergroup, Collodictyon triciliatum and Diphylleia rotans. From single cells, we generated high-quality draft genome assemblies estimated to be 70-84% complete, demonstrating the potential of long-read single-cell genomics to unlock genome diversity from uncultivated microbial eukaryotes.
Acharya, D.; Vembar, S. S.
Show abstract
Epigenetic regulation is central to the developmental progression and pathogenicity of the unicellular eukaryotic parasite Plasmodium falciparum; yet, the contribution of DNA base modifications remains poorly understood. One such modification, 8-oxoguanine (8-oxoG), which was initially identified as an oxidative lesion and a marker of DNA damage, has since emerged as a transcriptional regulator in advanced eukaryotes. Given that P. falciparum encounters a highly oxidative environment in human blood, we investigated the potential gene regulatory role of 8-oxoG during its intra-erythrocytic developmental cycle (IDC). Using immunodetection assays, we first confirmed the presence of 8-oxoG in P. falciparum genomic DNA and observed a gradual increase in 8-oxoG abundance from ring to schizont stages. We then optimized oxidative DNA immunoprecipitation sequencing (OxiDIP-seq) for the highly AT-rich parasite genome and generated genome-wide 8-oxoG profiles across four IDC timepoints, which revealed reproducible enrichment of 8-oxoG at discrete genomic loci, with more than 50% of the peaks stable across developmental stages. Notably, 8-oxoG accumulated at putative G-quadruplex-forming sequences in the parasite genome and preferentially localized within exonic regions of protein-coding genes, exhibiting a marked enrichment near STOP codons and within 3' untranslated regions. This in turn correlated with significantly higher steady-state transcript levels of 8-oxoG-marked genes, with stage-specific changes in 8-oxoG enrichment closely matching transcriptional activity. Furthermore, 8-oxoG-marked loci were preferentially associated with active and poised histone post-translational modifications, while showing no evidence of altered nucleosome occupancy. Collectively, these findings demonstrate that 8-oxoG is a widespread and non-random DNA modification in P. falciparum and suggest that it may function as an epigenetic mark associated with transcriptionally permissive chromatin and gene activation during parasite blood-stage development.